Papers by Alexander Miserlis Hoyle
Are Neural Topic Models Broken? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm . |
| Approach: | They show that neural topic models fare worse in both respects compared to an established classical method. |
| Outcome: | The proposed method outperforms the members of the ensemble in both respects. |
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? (2026.eacl-long)
Copied to clipboard
Jingwei Ni, Yu Fan, Vilém Zouhar, Donya Rooein, Alexander Miserlis Hoyle, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Elliott Ash
| Challenge: | Variation in human annotation (i.e., disagreements) is common in NLP, but it is unclear whether it is possible to model this variation in LLMs. |
| Approach: | They evaluate the influence of different reasoning settings on LLM disagreement modeling . RLVR-style reasoning degrades performance in disagreement modeling, they find . |
| Outcome: | The proposed reasoning settings improve LLM disagreement modeling, while RLVR-style reasoning degrades it. |
Improving Neural Topic Models using Knowledge Distillation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Current paradigms for transfer learning use general knowledge as a foundation for more specialized endeavors. |
| Approach: | They propose to combine probabilistic topic models and pretrained transformers to improve topic quality by using knowledge distillation. |
| Outcome: | The proposed framework improves topic quality over all estimated topics and in head-to-head comparisons of aligned topics. |
Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic Models (2025.acl-long)
Copied to clipboard
Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan Lee Boyd-Graber
| Challenge: | a common use of NLP is to facilitate the understanding of large document collections. |
| Approach: | They propose to use large language models to replace probabilistic topic models in real-world applications. |
| Outcome: | The proposed model generates more human-readable topics and shows higher average win probabilities than traditional models for data exploration. |
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure (2025.emnlp-main)
Copied to clipboard
| Challenge: | Embedding-based similarity metrics can be influenced by content dimensions and spurious attributes like the text’s source or language. |
| Approach: | They propose a debiasing algorithm that removes observed confounders from encoder representations and removes them from the encoder. |
| Outcome: | The proposed method improves on out-of-distribution benchmarks and on benchmarks, but performance is not affected. |
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification (2025.emnlp-demos)
Copied to clipboard
Chenfei Xiong, Jingwei Ni, Yu Fan, Vilém Zouhar, Donya Rooein, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Zhijing Jin, Mrinmaya Sachan, Markus Leippold, Dirk Hovy, Mennatallah El-Assady, Elliott Ash
| Challenge: | Social scientists often need to develop codebooks that can be reliable but require significant human effort. |
| Approach: | They propose a mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language models. |
| Outcome: | The proposed framework integrates human expertise with automatic annotation guided by large language models. |
How Persuasive Is Your Context? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Empirically, through aseries of experiments, we show that TPS captures a more nuanced notion of persuasiveness than previously proposed metrics. |
| Approach: | They introduce a targeted persuasion score to quantify how persuasive a given context is to an LM. |
| Outcome: | Empirically, the proposed model captures a more nuanced notion of persuasiveness than previously proposed metrics. |
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering (2025.acl-long)
Copied to clipboard
| Challenge: | Topic models and document clustering evaluations often use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. |
| Approach: | They propose a protocol for evaluating topic models and document clustering evaluations that uses crowdworker annotations to validate automated proxies. |
| Outcome: | The proposed protocol is scalable and easy to adapt to an LLM prompt. |
PairScale: Analyzing Attitude Change with Pairwise Comparisons (2025.findings-naacl)
Copied to clipboard
| Challenge: | a text-based framework for measuring attitudes in communities is proposed . the framework uses both implicit and explicit evidence in language to characterize attitudes . |
| Approach: | They propose a text-based framework for measuring attitudes in communities toward issues of interest using language. |
| Outcome: | The proposed framework is validated by examining attitudes on two high-profile issues in the u.s. |
Promoting Graph Awareness in Linearized Graph-to-Text Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent applications of pretrained transformers to linearizations of graph inputs yield stateof-the-art results on graph-to-text tasks. |
| Approach: | They propose to use pretrained transformers to encode local graph structures . they find they can improve the quality of models' implicit graph encodings . |
| Outcome: | The proposed models can encode local graph structures and reconstruct corrupted inputs. |
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments (2026.acl-long)
Copied to clipboard
Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag
| Challenge: | Apertus is a fully open suite of large language models (LLMs) designed to address responsibility shortcomings in today’s open model ecosystem, namely data responsibility and global representation. |
| Approach: | They propose to release a fully open suite of large language models (LLMs) that address data responsibility and global representation shortcomings in today’s open model ecosystem. |
| Outcome: | The proposed model is pretrained on openly available data and suppresses verbatim recall of data while retaining task performance. |
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling (P19-1)
Copied to clipboard
| Challenge: | a recent study has focused on the ways in which language is gendered . positive adjectives used to describe women are more often related to their bodies . |
| Approach: | They propose a model that models adjective choice and its sentiment given the natural gender of a head noun. |
| Outcome: | The proposed model shows that positive adjectives used to describe women are more often related to their bodies than positive adjective words used to explain men. |
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards? (2021.acl-long)
Copied to clipboard
| Challenge: | Rather than replacing leaderboards, we advocate a re-imagining of the model to highlight if and where progress is made. |
| Approach: | They propose a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. |
| Outcome: | The proposed model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. |
Combining Sentiment Lexica with a Multi-View Variational Autoencoder (N19-1)
Copied to clipboard
| Challenge: | a new model of sentiment lexica is being developed to combine disparate scales into a common representation. |
| Approach: | They propose a model that unifies disparate scales into a common latent representation . they evaluate a text classification task using nine English-Language sentiment datasets . |
| Outcome: | The proposed model outperforms six individual sentiment lexica and a simple combination thereof. |
Measuring scalar constructs in social science with LLMs (2025.emnlp-main)
Copied to clipboard
Hauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle
| Challenge: | Valid scalar measurement of skalar constructs is a fundamental task in text analysis. |
| Approach: | They evaluate four approaches to measuring scalar constructs using large language models . pairwise comparisons produced better measurements than prompting LLMs, they say . validation of skalar measurement enables wide range of substantive applications in social science research . |
| Outcome: | The proposed methods improve on pairwise comparisons and finetuning . the proposed methods can be used in social science research . |